Company Description
Kotak Securities Limited is one of India's leading financial services organizations and a subsidiary of Kotak Mahindra Bank. The company provides investment, trading, and wealth management solutions to clients across India. With a strong focus on technology, innovation, and customer experience, Kotak Securities continues to invest in scalable digital platforms and modern engineering capabilities. The organization offers long-term career opportunities, structured growth, and exposure to large-scale technology systems within a regulated financial services environment.
Role Description
The Lead, DevOps and Platform Engineering will lead the development and management of core infrastructure, cloud platforms, and reliability engineering capabilities at Kotak Securities. This full-time role is based in Mumbai or Bengaluru and reports to the Head of SRE. The ideal candidate will have 10–15 years of experience in infrastructure, DevOps, platform engineering, or site reliability engineering, including strong technical leadership experience. Key responsibilities include designing and managing Kubernetes and virtual machine infrastructure, implementing infrastructure as code, building scalable CI/CD and GitOps pipelines, and establishing observability, reliability, and security standards. The role will also lead platform engineering initiatives supporting the modernization of trading systems, datacenter migration, cloud infrastructure, and critical financial applications. The lead will be responsible for platform performance, incident management, infrastructure security, regulatory compliance, capacity planning, and cost optimization while building and managing a high-performing platform engineering team.
Qualifications
- 10–15 years of experience in infrastructure, DevOps, platform engineering, or site reliability engineering, including at least 3 years of engineering leadership experience.
- Deep hands-on expertise in production Kubernetes environments, including cluster architecture, control plane operations, upgrades, networking, service mesh, autoscaling, and resource management.
- Strong experience designing and managing cloud and on-premise infrastructure using Terraform or equivalent infrastructure-as-code tools.
- Expertise in AWS or comparable cloud platforms, including VPC architecture, subnet design, private connectivity, load balancing, DNS, IAM, and cloud security.
- Proven experience building and managing CI/CD pipelines, GitOps workflows, and release engineering using tools such as ArgoCD, Flux, or equivalent.
- Strong understanding of Linux systems, kernel tuning, networking, TCP/IP, filesystem performance, and troubleshooting using tools such as perf and eBPF.
- Hands-on programming experience in Go, Python, or equivalent languages for developing automation, platform tooling, Kubernetes operators, or controllers.
- Experience designing container infrastructure, including image hardening, vulnerability scanning, artifact signing, registry management, and software supply chain security.
- Expertise in observability and monitoring platforms, including OpenTelemetry, distributed tracing, metrics, logging pipelines, APM, and incident management integrations.
- Strong understanding of site reliability engineering practices, including service level objectives (SLOs), error budgets, incident response, postmortems, and change management.
- Experience implementing progressive delivery, automated rollbacks, infrastructure automation, and deployment strategies for mission-critical systems.
- Proven experience in infrastructure capacity planning, performance engineering, load testing, and optimization of latency-sensitive applications.
- Experience leading datacenter or cloud migration initiatives, including hybrid connectivity, network architecture, environment parity, cutover planning, and disaster recovery.
- Strong knowledge of infrastructure security, including secrets management, PKI, certificate lifecycle, workload identity, privileged access management, and policy-as-code frameworks such as OPA or Sentinel.
- Experience managing cloud infrastructure costs, including resource rightsizing, capacity optimization, commitment strategies, and cost attribution.
- Experience operating highly available, business-critical systems where downtime has significant commercial or regulatory consequences.
- Strong understanding of regulatory compliance, audit requirements, change control processes, and segregation of duties in enterprise infrastructure environments.
- Prior experience in capital markets, banking, payments, trading platforms, or another regulated financial services environment is preferred.
- Knowledge of low-latency trading infrastructure, exchange connectivity, market data systems, colocation environments, and PTP/NTP clock synchronization is advantageous.
- Experience with datacenter networking technologies, including BGP, hybrid routing, and firewall infrastructure migration, is beneficial.
- Operational experience managing databases and messaging systems such as PostgreSQL, Redis, Kafka, or MongoDB, including replication, failover, and backup verification, is preferred.
- Familiarity with SEBI CSCRF, RBI cybersecurity guidelines, or comparable regulatory frameworks is advantageous.
- Experience building a platform engineering or observability function from the ground up is highly desirable.
- Strong technical leadership, architectural decision-making, troubleshooting, stakeholder management, and cross-functional collaboration skills.